Xiaomi Open Sources Industrial-Level Target Speaker Speech Recognition Large Model CocktailASR-1
Xiaomi open sources the industrial-level target speaker speech recognition large model CocktailASR-1, aiming to solve the cocktail party problem where multiple people speak simultaneously. The model changes the traditional ASR input method, accurately locking onto and recognizing the target speaker's voice in noisy environments, improving recognition capabilities in complex scenarios, and providing new solutions for the industry.